Papers by Ahnaf Mozib Samin
ColorFoil: Investigating Color Blindness in Large Vision and Language Models (2025.naacl-srw)
Copied to clipboard
| Challenge: | Several studies indicate a lack of robustness of the models when dealing with complex linguistics and visual attributes. |
| Approach: | They propose a new V&L benchmark by creating color-related foils to assess the models’ perception ability to detect colors like red, white, green, etc. |
| Outcome: | The proposed benchmark evaluates seven state-of-the-art V&L models including CLIP, ViLT, GroupViT, and BridgeTower in a zero-shot setting and demonstrates that they have better color perception capabilities than CLIP and its variants and GroupVit. |